Papers with Multimodal information extraction

2 papers
MatViX: Multimodal Information Extraction from Visually Rich Articles (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods for multimodal information extraction are limited due to the multimodal nature of scientific articles and complex interconnections between data points.
Approach: They propose a benchmark to extract structured information from scientific articles . they use curated JSON files extracted from text, tables, and figures .
Outcome: The proposed benchmark is based on 324 full-length research articles and 1,688 complex structured JSON files curated by experts in polymer nanocomposites and biodegradation.
MCIL: Multimodal Counterfactual Instance Learning for Low-resource Entity-based Multimodal Information Extraction (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to perform multimodal information extraction only investigated entity-based tasks under supervised learning with adequate labeled data.
Approach: They propose to investigate the entity-based MIE tasks under the low-resource settings by decomposing the features into image, entity, and context factors.
Outcome: The proposed method is able to perform on two public MIE benchmark datasets and the experimental results confirm it.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations